Papers by Amalie Brogaard Pauli
Mind the Style Gap: Meta-Evaluation of Style and Attribute Transfer Metrics (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Large language models (LLMs) make it easy to rewrite a text in any style, but they are not straightforward when evaluating content preservation. |
| Approach: | They propose a large meta-evaluation of metrics for evaluating style and attribute transfer, focusing on content preservation. |
| Outcome: | The proposed method achieves higher alignment with human judgements than prompting a model of a similar size as an autorater. |
DaNE: A Named Entity Resource for Danish (2020.lrec-1)
Copied to clipboard
Rasmus Hvingelby, Amalie Brogaard Pauli, Maria Barrett, Christina Rosted, Lasse Malm Lidegaard, Anders Søgaard
| Challenge: | a named entity annotation for the Danish Universal Dependencies treebank is the largest publicly available named entity gold annotation. |
| Approach: | They propose a named entity annotation for the Danish Universal Dependencies treebank using the CoNLL-2003 annotation scheme DaNE. |
| Outcome: | The proposed annotations improve Danish named entity recognition over a recent cross-lingual approach and over norwegian training set. |
Measuring and Benchmarking Large Language Models’ Capabilities to Generate Persuasive Language (2025.naacl-long)
Copied to clipboard
| Challenge: | Recent studies have focused on specific domains or types of persuasion, but a general study has focused on how LLMs produce persuasive text. |
| Approach: | They construct a dataset to measure and benchmark the ability of Large Language Models (LLMs) to produce persuasive text. |
| Outcome: | The proposed model can be used to generate persuasive text across domains and domains. |
Can Humans Identify Domains? (2024.lrec-main)
Copied to clipboard
Maria Barrett, Max Müller-Eberstein, Elisa Bassignana, Amalie Brogaard Pauli, Mike Zhang, Rob van der Goot
| Challenge: | Textual domain is a crucial property within the Natural Language Processing community due to its effects on downstream model performance. |
| Approach: | They examine the level of human disagreement and the relative difficulty of each annotation task by training classifiers to perform the same task. |
| Outcome: | The authors show that human proficiency in identifying related intrinsic textual properties is low and that disagreements are high. |
Analysing Differences in Persuasive Language in LLM-Generated Text: Uncovering Stereotypical Gender Patterns (2026.findings-acl)
Copied to clipboard
| Challenge: | Prior work has shown that large language models can successfully persuade humans and amplify persuasive language. |
| Approach: | They propose a framework for evaluating how persuasive language generation is affected by recipient gender, sender intent, or output language. |
| Outcome: | The proposed framework varies persuasive language when the recipient gender is specified or when the sender intent is specified. |